Medical Image Analysis
○ Elsevier BV
Preprints posted in the last 90 days, ranked by how well they match Medical Image Analysis's content profile, based on 35 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit.
Chen, J.; Pham, T.-H.; Zhang, P.; Varghese, J.
Show abstract
Accurate measurement of intra-cardiac blood oxygen (O2) saturation is essential for cardiovascular assessment, yet current methods require invasive catheterization. T2-based cardiac magnetic resonance imaging (CMRI) enables non-invasive O2 quantification, but deep learning automation is constrained by scarce annotated data. We propose a unified self-supervised learning (SSL) framework integrating cine CMRI and T2 oximetry CMRI to learn generalizable representations without labels. Our approach pre-trains ResNet and vision transformer encoders using contrastive learning and masked image modeling on over 48,000 cardiac images. Pre-trained encoders are fine-tuned for O2 saturation regression with uncertainty quantification to enhance clinical trustworthiness. Our SSL framework significantly outperforms traditional radiomics and supervised baselines, with SimCLR pre-trained ResNet achieving a mean absolute error of 3.70, representing over 15\% improvement. These findings demonstrate SSL's potential to address annotation bottlenecks in non-invasive cardiac diagnostics.
Sriram, R.; Nenadic, I.; Shahrabani, E.; Goonewardena, S.; Yao, S.; Farrell, B.; Loring, Z.; Murthy, V. L.
Show abstract
We conducted a scaling evaluation of unlabeled pretraining for electrocardiogram foundation model performance. One-dimensional vision transformer masked autoencoders were pretrained across increasing ECG volumes and fine-tuned for rhythm, morphology, diagnostic, and structural heart disease tasks. Models pretrained below 400,000 ECGs failed to consistently exceed controls without self-supervised pre-training, whereas 600,000 to 800,000 ECGs improved AUROC across tasks, suggesting a minimum threshold for effective ECG representation learning.
Yang, K.; Shi, P.; Huang, H.; Musio, F.; Baazaoui, H.; Aydin, O. U.; Hilbert, A.; Hamadache, R. E.; Yalcin, C.; Zhang, M.; Falcetta, D.; de la Rosa, E.; Shit, S.; Prabhakar, C.; Wittmann, B.; Rokuss, M. R.; Kirchhoff, Y.; Al-Maskari, R.; Hoeher, L.; Juchler, N.; Casamitjana, A.; Cleary, J.; Schmick, A.; Baumgartner, P.; Deseoe, J.; Vandans, O.; Lee, D.; Oh, K.; LaBella, D.; Mazher, M.; Niederer, S. A.; Qayyum, A.; Liu, Y.; Chen, J.; Kim, W.; Asawalertsak, N.; Kim, M.; Shin, D.; Park, S.-H.; Kikuchi, S.; Zhang, Y.; Liu, J.; Cui, Y.; Qiu, Y.; Verschuur, A.; Zhang, J.; van der Schaaf, I.; Su, R.;
Show abstract
We present the TopBrain 2025 Challenge, the first benchmark for fine-grained multiclass segmentation of the whole brain vasculature in both computed tomography angiography (CTA) and magnetic resonance angiography (MRA). Building on the TopCoW challenge, TopBrain scales vessel annotation from the Circle of Willis to the entire brain, introducing a dataset of 90 annotated volumes across 48 landmark vessel classes spanning arterial and venous systems, of which 50 training volumes are publicly released. Vessel definitions were consolidated from established neuroanatomical references into a unified annotation scheme, and vessel caliber measurements along the centerline are reported for the first time across the whole brain vascular anatomy. To address the unique challenges of multiclass brain vessel segmentation, we propose an evaluation framework that accounts for detection in segmentation performance, assesses anatomical plausibility, and introduces novel contamination metrics that characterize inter-class prediction errors. Fifteen teams from over 220 registered participants submitted algorithms to the benchmark. The top-performing teams built on nnUNet with principled system design choices, achieving around 80% Dice scores, near-zero invalid neighbor counts, over 60% F1 scores for side-road vessels, and below 18% foreground contamination ratio. Larger vessels are easier to segment, while smaller and more complex vessels remain the true bottleneck. The annotated datasets and podium-finish algorithms are made publicly available on Zenodo.
Donle, L.; Phillips, M.; Gaber, F.; Ramesh, S.; Sacco, M.; Hautaniemi, S.; Virtanen, A.; Bressem, K.; Adams, L.; Goon, K.; Nevins, E.; Robinett, R. A.; Kochanny, S.; Hassan, S.; Dolezal, J.; Pearson, A. T.; Lengyel, E.
Show abstract
Medical foundation models compress biomedical data into embeddings that support diverse downstream clinical tasks. However, successful model deployment is hampered by performance degradation on external data. It is recognized that embeddings capture acquisition signatures, such as hardware and technical differences, in addition to biology. Effective harmonization must remove the acquisition signature while preserving biological signals, a trade-off that current methods fail to balance adequately. Input-level normalization fails to eliminate acquisition signatures from embeddings, whereas embedding-level methods adjust features in an untargeted manner. We present FEATMAP, a harmonization approach that models acquisition signatures as geometric distortions between manifolds of similarly arranged embeddings. Using paired data that isolate the effect of acquisition signatures, FEATMAP fits a single global affine transformation per foundation model to correct acquisition signatures directly in the embedding space. This targeted, reusable correction aims to preserve biological and demographic variation while harmonizing across acquisition signatures. Across scanner and foundation-model harmonization in digital pathology and field-strength harmonization in brain MRI, FEATMAP improves cross-condition embedding similarity, reduces performance gaps without retraining, and suggests potential for the alignment of disparate embedding spaces.
You, L.; Dang, H.; Wang, H.; Matta, E.; zhou, X.
Show abstract
Image-based liver Couinaud segmentation is designed to automatically provide the locations of suspicious objects in liver CT/MR images. Once achieved, the physicians will be guided to the target slice and area where the suspicious node is located. However, conventional algorithms trained primarily on healthy liver images often fail to generalize to Hepatocellular Carcinoma (HCC) cases due to pathological structural distortions. In this work, we propose a robust two-stage framework that integrates a 3D Unet with a 3D Anatomical Structure-Guided Graph Convolutional Network (3D GCN). This two-stage strategy effectively isolates the liver volume to eliminate structural noise from neighboring organs, such as the spleen, allowing the framework to focus exclusively on the complex 3D anatomical relationships among the eight segments. To ensure the topological consistency required for global spatial reasoning, we implement a standardized preprocessing pipeline that normalizes liver-only volumes to exactly 50 frames along the z-axis. By combining a lightweight 3D UNet backbone with the 3D GCN for refined boundary reasoning, our model demonstrates superior generalization performance on unseen clinical datasets, achieving a mean Dice score of 0.828 in blind testing. By releasing our code and pretrained weights, we aim to provide the first publicly available deep learning resource for robust Couinaud segmentation.
Encin, A.; Pepe, I. G.; Chatelain, Y.; Dickie, E.; Glatard, T.
Show abstract
We demonstrate that features extracted from structural MRI using un-CNN, an untrained convolutional neural network, achieve predictive performance comparable to or exceeding that of state-of-the-art pretrained foundation models across three structural MRI datasets and three downstream tasks. Un-CNN extends a classical 3D CNN architecture with multi-channel inputs, a hierarchical encoder with multi-scale feature aggregation, and covariance pooling. Untrained CNNs circumvent several key limitations of trained models, including high computational cost and memory requirements, the need to distribute large model weights, risks of data leakage, and challenges in reproducibility.
Nazir, A.; Cheema, M. N.; Hsu, Y.-C.; Jiang, X.; Zhu, J.-J.; Harmanci, A. S.; Harmanci, A.
Show abstract
Gliomas are aggressive primary brain tumors that necessitate critical molecular biomarker predictions for optimal clinical decision-making. Traditional assessment relies on surgical tumor specimens analysis, which carries procedural risks and sampling bias due to tumor heterogeneity. Existing deep learning methods for non-invasive prediction lack real-time applicability, remain resource-intensive, and are frequently trained on narrowly represented datasets. We present GlioVision, a framework built on the MONAI library to process multimodal data, including glioma MRI and molecular labels, to predict and identify, non-invasively, four major glioma molecular biomarkers: IDH mutation, 1p/19q co-deletion, MGMT methylation, and WHO grade. The core architecture comprises Spatially and Channel-wise Recalibrated 3D DenseNet (SCRU-DenseNet), which utilizes a computational attention gate and an Adaptive Contrast-Specific Processing Stream (ACPS) to tackle multi-site, heterogeneous datasets. We introduced the Confidence-Filtered Predictive Manifold (CFPM) to manage uncertainty by excluding predictions with low confidence. GlioVision is trained and validated on the largest multi-cohort datasets, achieving strong biomarker prediction with AUCs of (IDH 0.94, 1p/19q 0.87, MGMT 0.86, WHO grades 0.92), supporting molecularly defined glioma diagnosis under the WHO 2021 classification guidelines. Finally, we provide a Differential Training Integrity Assessment (DTI-A) to analyze routes of MRI data privacy protections through model obfuscation. Our results advance the codebase, model release, and leakage considerations around MRI data analysis literature.
Tustison, N. J.; Avants, B. B.; Cook, P. A.; Gee, J. C.; Stone, J. R.
Show abstract
In modeling complex probability distributions, normalizing flows provide exact-likelihood, bijective mappings between empirical data and tractable latent spaces. Building on this foundation, latent-aligned multiview normalizing (LAMNr) flows leverage these salient properties to learn shared latent subspaces across heterogeneous, multimodal datasets while simultaneously topologically unfolding the sampled data manifold into a continuous vector space. Formal latent-alignment constraints are used to model shared structural features separate from view-specific variations, coordinating latent projections into a shared geometric subspace. By applying this transformation in the context of biological imaging, the framework establishes a potential basis for a deep learning interpretation of foundational computational anatomy concepts, such as the population template, latent distances, and geodesic pairwise image interpolation. Additionally, the proposed framework enables closed-form conditional modeling for exact cross-view imputation and other latent space manipulations. Evaluations and illustrations on both imaging-derived phenotypes (IDPs) and multimodal MRI demonstrate the proposed framework and potential applications. To further motivate our work, we provide a robust and comprehensive, 2D and 3D open-source implementation in PyTorch, natively integrated with the ANTsX ecosystem (i.e., ANTsTorch) for efficient training and subsequent data transformation, manipulation, and analysis.
Wen, R.; Zhang, J.; Liang, Z.
Show abstract
Diffusion MRI (dMRI) tractography provides a non-invasive method for mapping whole-brain structural connectivity. However, its application is limited by substantial false-positive and false-negative connections. While deep learning based methods have shown promise in improving tractography, most rely on training data derived from conventional dMRI tractography, therefore inheriting the same limitations. Here, we introduce FiberLM, an attention-based Transformer model for mouse brain tractography. The model was trained using a whole-brain streamline dataset based on viral tracer data from the Allen Mouse Brain Connectivity Atlas (AMBCA), allowing the model to learn the properties of both local and long-range axonal trajectories through self-attention. FiberLM was applied to predict anatomically plausible axonal trajectories from ex vivo high-resolution mouse brain dMRI data. Quantitative evaluations demonstrated that FiberLM significantly reduced false-positive and false-negative connections, improved spatial agreement with tracer-defined pathways, and generated whole-brain connectomes that more closely approximated AMBCA results compared to conventional tractography. These findings suggest FiberLM as a potential tool for accurate reconstruction of mouse brain structural connectomics.
Bors, S.; Beyeler, M.; Trofimova, O.; VascX Consortium, ; Presby, D.; Bontempi, D.; Bergmann, S.
Show abstract
Deep learning models based on Vision Transformers (ViTs) have shown strong performance in retinal fundus imaging, but their interpretability remains poorly understood. In particular, attention-based attribution methods are widely used to explain ViT predictions, despite limited evaluation of their faithfulness and biological relevance in medical imaging. Here, we systematically benchmark four attention-based interpretability methods for RETFound, a retinal ViT-based foundation model, that we previously fine-tuned to predict 17 retinal vascular phenotypes from UK Biobank fundus images1. We compare raw attention, attention rollout, gradient-weighted attention rollout, and Chefers hybrid relevance-based method using both qualitative visualisation and quantitative evaluation frameworks. To assess attribution faithfulness, we perform perturbation-based deletion and insertion experiments, quantifying changes in model predictions as highly attended image regions are progressively removed or restored. To evaluate biological specificity, we run structure-aware analyses combining attribution maps with vessel segmentation and artery-vein labels through the Relative ratio of Attention Intensity (RAI) metric. Across models, attribution maps differed substantially depending on the selected interpretability method, highlighting the need for rigorous quantitative evaluation. Among the evaluated approaches, gradient-weighted attention rollout consistently achieved the strongest perturbation performance and produced attribution maps most closely aligned with the anatomical definition of the predicted retinal traits. Furthermore, vessel-type specific models systematically concentrate attention on the corresponding vascular structures despite being trained using only a single scalar value per image as supervision. These findings demonstrate that attention-based attribution methods capture biologically meaningful vascular representations, while also revealing method-dependent variability in attribution behaviour. This work provides a quantitative framework for evaluating interpretability methods in medical imaging with annotated segmentation and contributes toward more transparent and biologically grounded medical AI systems.
Kumada, C.; Hiroyasu, T.; Hiwa, S.
Show abstract
Structural connectivity (SC) data are crucial for brain network analysis, but SC-based machine learning often suffers from limited data availability, hindering model generalization and robustness. Although data augmentation using deep generative models has attracted increasing attention, it remains unclear how different models capture the complex topological features of SC data. To clarify the learning characteristics of deep generative models for SC generation, this study compares three representative models: variational autoencoder (VAE), Wasserstein GAN with gradient penalty (WGAN-GP), and denoising diffusion probabilistic models (DDPM). We systematically evaluated these models using both synthetic datasets with known characteristics and real-world SC data. Generation quality was assessed using graph-theoretic metric comparisons and visual inspection of the generated adjacency matrices. WGAN-GP showed relatively stable performance across datasets and metrics, without severe performance degradation across evaluation settings. In contrast, VAE and DDPM performed well in specific aspects but were more sensitive to data characteristics. These findings suggest that WGAN-GP may serve as the most balanced baseline for future SC data augmentation studies, whereas VAE and DDPM may be useful depending on the target application and structural properties of interest. Furthermore, because all models struggled to fully reproduce strict global constraints such as planarity, our results suggest that standard generative models may be insufficient to capture the complex topological features of SC data. This highlights the importance of incorporating the desired structural properties into the training or generation process.
Cajas, S.; Marzullo, A.; Kapadia, S.; Santos, F.; Ocampo Osorio, F.; Kong, Q.; Quarta, A.; Kuo, P.-C.; Patel, M.; Rojas Sillery, R. I.; Celi, L. A.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWShortcut learning poses a significant challenge in clinical artificial intelligence, as models may rely on spurious signals rather than clinically relevant features, leading to biased predictions and poor generalization. Existing detection methods are fragmented and lack systematic evaluation across datasets and model architectures. To address this issue, we propose ShortKit-ML, an open-source Python framework for unified shortcut analysis in embedding spaces. The framework integrates over 20 detection methods and six mitigation strategies within a modular pipeline, encompassing embedding analysis, fairness metrics, training dynamics, causal methods, explainability, and representation analysis. We evaluate the framework on chest X-ray datasets (CheXpert and MIMIC-CXR), synthetic benchmarks, and an out-of-domain dataset (CelebA). Experimental results demonstrate that multi-method auditing provides more stable and interpretable evidence than individual methods, while detector disagreement reveals meaningful representational differences. The proposed framework offers automated reporting, interactive visualization, and is available as a pip-installable package. The source code and documentation are publicly available at https://github.com/criticaldata/ShortKit-ML and https://criticaldata.github.io/ShortKit-ML/.
Zhu, S.; Dinsdale, N. K.; Jbabdi, S.; Miller, K.; Howard, A.
Show abstract
Characterising human brain connectivity remains a major challenge in neuroscience. Multimodal datasets combining diffusion MRI with high-resolution microscopy in the same brain offer a unique link between macroscopic imaging and microstructural detail, but we lack tools to leverage these data to improve connectivity estimates for in vivo human imaging. We present a deep learning model that predicts high-resolution microscopy-informed fibre orientations from diffusion MRI. The model uses microscopy-derived three-dimensional fibre orientation maps as biologically grounded training targets. It is trained on a bespoke macaque dataset integrating in vivo MRI, postmortem MRI, and whole-brain microscopy, and then translated to in vivo human imaging. We use domain adaptation to predict fibre orientations from diverse MRI datasets: first to bridge differences in tissue state in the macaque (postmortem to in vivo), and then to generalise across species (macaque to human). Our method derives microscale-informed fibre architecture from diffusion MRI without requiring microscopy at inference. It leverages data that can easily be acquired only in animal models whilst generalising to in vivo human diffusion MRI with minimal acquisition requirements. The microscopy-informed fibre orientation distributions support biologically meaningful tractography, enhancing superficial white matter and cortical-subcortical pathway delineation for in vivo human data. More broadly, this work establishes a general framework for transferring microstructural information from microscopy to non-invasive imaging, enabling biologically informed mapping of brain connectivity.
Matsulevits, A.; Koch, A.; Mahe-Verdure, C.; Bendszus, M.; Hilbert, A.; Boullet, M.; Marnat, G.; Mutke, M.; Aydin, O.; Olindo, S.; Sibon, I.; Frey, D.; Thiebaut de Schotten, M.; Tourdias, T.
Show abstract
BackgroundMagnetic resonance imaging (MRI) is critical for acute stroke triage, but time-consuming, and often requires contrast injection for perfusion imaging. This study aimed to synthesize T-map perfusion maps from routinely available, non-contrast DWI and FLAIR using deep generative models. We hypothesized that relevant perfusion information could be inferred from these modalities to streamline imaging and reduce reliance on dynamic susceptibility contrast perfusion. MethodsAcute MRI data from 355 patients with anterior circulation stroke, including dynamic susceptibility contrast perfusion, were retrospectively collected from two European centers (Heidelberg: 2010-2018; Bordeaux: 2021-2022). Six versions of a denoising diffusion probabilistic model (DDPM) and a GAN architecture were trained to generate synthetic T-max perfusion maps from DWI, FLAIR, and infarct core mask as inputs. Performance was assessed by comparing synthetic and ground truth T-max maps using image similarity metrics. Regions with T-max >6s were compared using Dice coefficients, and mismatch volume distributions were analyzed. An ablation study quantified the contribution of each input. ResultsThe best performance was achieved by a DDPM with a 2.5D architecture using DWI, FLAIR, infarct core mask, and a perfusion-weighted loss function. It produced synthetic perfusion T-max maps with high similarity to ground truth under 110 seconds. The model showed strong spatial overlap for T-max >6s regions in internal validation (average Dice = 0.82, SD = 0.08), and external validation average (Dice 0.59, SD = 0.13), respectively. Synthetic maps closely matched ground-truth mismatch distributions, capturing key perfusion patterns. The infarct core mask played a critical role in model performance, alongside DWI and FLAIR inputs. ConclusionsWe propose a non-invasive, scalable framework to generate synthetic T-max perfusion maps from non-contrast MRI. This approach could expand access to perfusion data in acute stroke, shorten imaging protocols, and accelerate treatment decisions by eliminating the need for contrast-enhanced acquisition. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=200 SRC="FIGDIR/small/684079v2_ufig1.gif" ALT="Figure 1"> View larger version (94K): org.highwire.dtl.DTLVardef@164235forg.highwire.dtl.DTLVardef@14e5489org.highwire.dtl.DTLVardef@190214eorg.highwire.dtl.DTLVardef@17a9e3a_HPS_FORMAT_FIGEXP M_FIG C_FIG
Vlachas, P.; Nonchev, K.; Koelzer, V.; Ratsch, G.
Show abstract
Spatial transcriptomics couples hematoxylin and eosin (H&E) tissue morphology with spatially resolved gene expression (GE). However, generative models that exploit this coupling to synthesize tissue images from transcriptomic profiles remain scarce. We present STMDiT (Spatial Transcriptomics and Morphology Diffusion Transformer), a diffusion transformer that synthesizes H&E histopathology patches conditioned jointly on morphological embeddings and transcriptomic profiles. Building on PixCell (Yellapragada et al., 2025), we integrate gene expression from a frozen CancerFoundation encoder (Theus et al., 2024) through adaptive layer normalization and per-block cross-attention, and we train under dual classifier-free guidance with independent modality dropout. On the 10x TuPro Visium melanoma cohort, GE conditioning improves both image quality over the no-GE PixCell-B baseline (best FID = 252.9 vs 330.7) and transcriptomic fidelity (best AUC = 0.267 vs 0.229, reaching 82% of the real-tile ceiling). Training with DeepSpots predicted-transcriptomics pseudo-labels (PTPL) uniquely transfers zero-shot to TCGA SKCM, an out-of-distribution (OOD) H&E-only melanoma cohort: PTPL-XAttn-PMA-B reaches FID = 690.0, a 57-point improvement over the no-GE baseline (747.1), with a within-model GE-ablation effect of {Delta}OOD = +309.5, enabling virtual tissue synthesis beyond native spatial-transcriptomics coverage. Our results indicate that gene-expression conditioning produces morphologically distinct tissue images and supports virtual tissue simulation for hypothesis testing in computational pathology.
Bit, S.; Guney, O. B.; Jia, S.; Kolachalama, V. B.
Show abstract
Automated interpretation of neuroimaging studies requires simultaneous assessment of multiple imaging evidence variables, each tied to distinct anatomical structures. Vision-language models (VLMs) offer a unified framework for multi-task analysis, but adapting pre-trained VLMs remains challenging. Full fine-tuning is computationally prohibitive, and joint multi-task training requires simultaneous access to all task data, which is often infeasible in clinical settings. Although model merging enables multi-task composition without joint re-training, existing methods focus on post-hoc algorithms with limited extension to VLMs and minimal application to neuroimaging. Here, we present GRadient-guided Adapter Merging (GRAM), a layer-selective low-rank adaptation (LoRA)-based fine-tuning and merging framework for multi-task neuroimaging visual question-answering (VQA). GRAM uses a gradient ratio that contrasts class-specific gradients to identify task-discriminative layers, and applies subspace-constrained projected gradient descent to restrict LoRA updates to directions consistent with the geometry of the pre-trained model. We leveraged a structured VQA benchmark, developed from the National Alzheimer's Coordinating Center (NACC) dataset, that pairs multi-sequence brain MRI studies with question-answer pairs across clinically relevant imaging evidence variables. Experiments on the VQA benchmark showed that GRAM outperformed or matched all-layer LoRA fine-tuning and a standard merging baseline while reducing inter-task interference during merging, and approached or surpassed the performance of joint multi-task training without joint re-training.
Majid, I.; Wang, M.
Show abstract
Purpose: To determine whether disease-aware adversarial perturbations can reduce demographic recoverability encoded in color fundus photographs (CFPs) while preserving glaucoma-related diagnostic features. Design: Retrospective analysis of a single-institution retinal imaging dataset using adversarial machine-learning experiments. Participants: A total of 4,271 patients contributing 13,959 CFPs from Massachusetts Eye and Ear. Methods: Vision Transformer (ViT) was trained for glaucoma detection and for prediction of race, sex, and ethnicity. Standard and disease-aware (DA) variants of four adversarial attacks--Fast Gradient Sign Method (FGSM), Projected Gradient Descent (PGD), Carlini & Wagner (C&W), and a diffusion-based attack--were applied to suppress demographic prediction; DA attacks augmented the adversarial objective with a disease-preservation term. Cross-architecture transferability was assessed by generating perturbations on ViT and applying them to ResNet50 and EfficientNetB0. Main Outcome Measures: Area under the receiver operating characteristic curve (AUC) and accuracy for glaucoma and demographic classification before and after perturbation, and disease-preservation and attack transferability across architectures. Results: At baseline, CFPs encoded both glaucoma-related and demographic information. Glaucoma detection AUCs were 0.958 (95% CI, 0.949-0.967), 0.960 (95% CI, 0.951-0.967), and 0.963 (95% CI, 0.955-0.971) in the race, sex, and ethnicity analysis cohorts, respectively. Demographic prediction performance was also high, with AUCs of 0.955 (95% CI, 0.945-0.963) for race, 0.983 (95% CI, 0.977-0.988) for sex, and 0.992 (95% CI, 0.987-0.996) for ethnicity. Standard attacks substantially reduced demographic AUC but often degraded glaucoma detection. Disease-aware optimization improved disease preservation while maintaining demographic suppression. Using a prespecified success criterion of at least 90% disease AUC preservation and demographic AUC reduction to 30% or less of baseline, DA-PGD and DA-Diffusion succeeded across race, sex, and ethnicity; DA-C&W succeeded for sex and ethnicity. Cross-architecture transferability experiments demonstrated that disease preservation transferred more robustly than demographic suppression. Conclusions: Disease-aware adversarial perturbations reduced the recoverability of demographic information in CFPs under white-box conditions while preserving glaucoma-relevant features, suggesting these representations are partially separable. Reduced demographic recoverability did not fully transfer across architectures, highlighting the need for architecture-agnostic methods.
Shenoy, A. R.; Mendez, T.
Show abstract
Stroke is a leading cause of death and long-term disability worldwide, affecting approximately 15 million individuals annually. Prompt and accurate subtype differentiation between ischemic and hemorrhagic stroke is clinically critical, as the two conditions demand diametrically opposite interventions - thrombolytic therapy versus surgical decompression. Yet the majority of existing deep learning approaches reduce this problem to binary detection, and virtually none address the opacity of their decision-making in a clinically actionable manner. We present CerebAI, an explainable, deployment-oriented three-class CT stroke classification system built on a fine-tuned ConvNeXt-Base backbone with Integrated Gradients (IG) attribution. Trained on 6,774 non-contrast CT scans stratified across No Stroke, Ischemic Stroke, and Hemorrhagic Stroke, CerebAI achieves a weighted F1-score of 0.9746 (95% CI: [0.9625, 0.9851]), accuracy of 97.47%, macro-averaged AUC of 0.9921, mean Intersection-over-Union (mIoU) of 0.9276, Expected Calibration Error (ECE) of 0.0115, mean Brier Score of 0.0150, and Cohen's {kappa} of 0.9483 - surpassing ResNet-50, EfficientNet-B4, and Vision Transformer (ViT-B/16) baselines across all reported metrics. Integrated Gradients produce pixel-precise saliency maps that localize pathological regions with greater anatomical fidelity than Gradient-weighted Class Activation Mapping (Grad-CAM), a finding we support with side-by-side qualitative comparison. CerebAI additionally incorporates a native DICOM processing pipeline to facilitate future clinical translation. Code and model weights are publicly available to support reproducibility and further research.
Dillon, T. M.; Quevedo Moreno, D.; Rutherford, E. K.; Ayers, B.; Salomon, B.; Kubi, B.; Thomas, J.; Roche, E.
Show abstract
Minimally invasive endovascular procedures offer reduced surgical trauma, shorter recovery times, and improved outcomes, but rely on 2D fluoroscopic X-ray imaging, which provides limited depth perception and exposes patients and clinicians to ionizing radiation. Here we present an augmented reality (AR) system that fuses intravascular ultrasound (IVUS) and electromagnetic (EM) position tracking with preoperative computed tomography (CT) to produce an anatomically accurate, deformation-corrected navigational reference. A robotic device performs ECG-gated pullback of the IVUS probe, capturing 4D aortic motion across the cardiac cycle. We introduce a deep learning architecture for extracting vascular lumen boundaries and side-branch orifices from artifact-prone IVUS streams, and a semantically driven non-rigid CT-IVUS fusion pipeline robust to false positive landmarks. We evaluate the platform with trained surgeons in benchtop phantom studies and in-vivo ovine models, and demonstrate its application to fenestrated endovascular aneurysm repair (FEVAR). Compared to fluoroscopy alone, AR guidance significantly reduces cannulation time, radiation exposure, and cognitive workload, while improving procedural efficiency and safety. Our IVUS-EM and CT aortic datasets are released open source.
Gao, Y.; Li, J.; Xu, J.; Li, Q.; Li, Z.; Shi, Y.; ZHao, G.; Wu, X.; Zhang, Y.
Show abstract
Accurate and robust classification of medical pathology images is pivotal for computer-aided diagnosis. However, the deployment of deep learning models in high-throughput clinical screening faces a fundamental challenge: the trade-off between diagnostic accuracy and computational efficiency. Current lightweight architectures, while reducing parameter complexity through grouped convolutions, often lead to cross-channel information isolation and diminished representational capacity. In this paper, we propose TetraFuse, a novel framework that systematically integrates features from four complementary domains: space, channel, statistics, and frequency. TetraFuse introduces a novel Cross-Channel Dynamic Aggregation (CCDA) paradigm that reconstructs global channel topology with negligible computational overhead, resolving the inter-group isolation issue. To balance perceptual fidelity and efficiency, we design a stage-aware local enhancement mechanism: Local Variance-Guided Enhancer (LVGE) is employed to filter out shallow-stage background noise, while High-Frequency Boundary Injection (HFBI) reinforces deep-stage pathological contours, preventing spatial over-smoothing. Experimental results on the COVID-19, ISIC 2018, and Kvasir datasets confirm that TetraFuse outperforms state-of-the-art (SOTA) methods. Notably, TetraFuse-Tiny achieves a transformative 91.53% reduction in FLOPs compared to ResNet50; on the Kvasir dataset, it achieved an accuracy of 0.926 and an AUC of 0.994 with only 0.345G FLOPs. By combining high representational power with minimal computational demand, TetraFuse offers a scalable solution for large-scale medical image analysis, especially in resource-constrained clinical environments.